Papers with Spatio-Temporal Video Question Answering
TVQA+: Spatio-Temporal Grounding for Video Question Answering (2020.acl-main)
Copied to clipboard
| Challenge: | Existing video QA datasets only contain QA pairs without labels for key clips or regions needed to answer the question. |
| Approach: | They propose a framework that grounds evidence in both spatial and temporal domains to answer questions about videos using bounding boxes. |
| Outcome: | The proposed framework can produce interpretable spatio-temporal attention visualizations. |